Washington Women’s Heritage Project (WWHP) Oral History Corpus
Methodology
Version 1.0
July 2026

PURPOSE OF THIS DOCUMENT

This document describes the provenance, selection, preparation,
normalization, and limitations of the Washington Women’s Heritage
Project Oral History Corpus.

It is intended to support transparency, reproducibility, and informed
reuse of the dataset for computational analysis.

CORPUS DESCRIPTION

The corpus consists of oral history transcripts originating from the
Washington Women’s Heritage Project (WWHP), a statewide initiative
documenting the lives and experiences of women in Washington State.

Corpus statistics:

- 48 interviews
- 228 text documents
- Approximately 1,010,420 words
- Interviews conducted between 1980 and 1998
- Geographic focus: Washington State

SOURCE COLLECTION

Source collection:

Washington Women’s Heritage Project records
University of Washington Libraries, Special Collections
Archives West identifier: ark:/80444/xv43023

Collection Guide:
https://archiveswest.orbiscascade.org/ark:80444/xv43023

The archival collection contains:

- Oral history audio recordings
- Transcript materials
- Project documentation
- Photographs
- Related project records

CORPUS SELECTION

This corpus represents a subset of the broader WWHP collection; transformation of the remaining oral histories is ongoing as of July 2026; subsequent transcripts will be added to this data set as numbered versions (2.0, 3.0, etc)

Selection criteria included:

-The University of Washington Libraries had the rights to make an interview available.

Interviews not meeting this criteria were not included.

Consequently, the corpus should not be regarded as complete or
statistically representative of the entire WWHP collection.

TRANSCRIPT PREPARATION

Original source materials consisted of audio cassette recordings. Cassettes were digitized; the digitized audio was uploaded into the transcription platform Otter.ai. Automated transcript was then edited while listening to the digitized recording; the speaker names were added in place of the automated output of Speaker 1, Speaker 2, etc. 
with oral history interviews. A transcript was created for each side of a cassette; if an interview was recorded on the A and B sides of two cassettes, 4 transcripts were generated. For this version 1.0 there are 48 interviews and 288 transcripts.

Transcripts were converted into plain text files and organized as
individual documents for computational analysis.

Output format:

- Plain text (.txt)
- UTF-8 encoding
- One document per transcripts format

DATA NORMALIZATION

The corpus was normalized to facilitate machine processing and
reproducible research.

Normalization activities included formatting transcripts into a
consistent machine-readable structure and standardizing text file
output.

The objective was to improve interoperability with common text-mining
and natural language processing workflows while preserving the
substantive textual content of the interviews.

METADATA CREATION

Metadata records were created and stored in metadata.csv.

Metadata fields include:

- id
- title
- interviewee
- interviewer
- date
- location
- collection
- source
- language
- url

These metadata fields enable filtering, grouping, and contextual
analysis of interview content.

FILE IDENTIFICATION

Each text file is assigned a unique identifier that corresponds to the
metadata record in metadata.csv.

Example:

wauar_3416-001-wwhp_alexanderLavinia01_side1.txt

Researchers should use the identifier field to link transcript files
with metadata records.

INTENDED USES

The corpus was developed to support:

- Text and data mining
- Natural language processing
- Concordance analysis
- Topic modeling
- Named entity recognition
- Comparative textual analysis
- Digital humanities research

RELATIONSHIP TO ORIGINAL MATERIALS

This corpus is a derivative dataset created from archival materials.

The dataset does not include:

- Audio recordings
- Full archival description
- Project documentation
- Other associated collection materials

Computational analysis should therefore be supplemented with
consultation of the original archival collection when historical
interpretation is important.

LIMITATIONS

Researchers should consider the following limitations:

1. Transcripted speech differs from recorded speech.

   The textual corpus does not capture tone, emphasis, pauses,
   inflection, or other vocal characteristics.

2. Non-verbal communication is absent.

   Gestures, facial expressions, and environmental context are not
   represented.

3. Archival context has been reduced.

   Collection-level and item-level contextual information available in
   the archival repository is not fully represented within the corpus.

4. Selection effects are present.

   Only interviews available and suitable for computational reuse were
   included.

RIGHTS AND REUSE

This dataset is distributed under:

Creative Commons Attribution-NonCommercial 4.0 International
(CC BY-NC 4.0)

Researchers are responsible for determining whether additional
restrictions apply to the underlying archival materials.

VERSION HISTORY

Version 1.0 — July 2026

Initial release of the WWHP Oral History Corpus with normalized
transcripts and structured metadata.